AAI 2026: AMD ROCm.ai Accelerates AI Development Across AMD Platforms
New AI-native developer experience enables the use of natural language and leading AI coding assistants to install, deploy, troubleshoot and optimize AI workloads on AMD platforms.

What’s the News? AMD announced ROCm.ai, an AI-native software experience for AI development across AMD platforms. It combines AI-assisted development, intelligent deployment and AI-powered software optimization into a unified experience that helps developers move quickly from intent to production. Introduced at Advancing AI 2026, ROCm.ai includes ROCm CLI – a new command-line experience for installing, validating, serving and managing AI workloads – and AMD Skills, bringing AMD-authored expertise into leading AI coding assistants such as Claude, Cursor and Codex. It also includes Hyperloom, a new open-source, agentic system aimed at automating the time-consuming task of optimizing end-to-end inference workloads. Together, ROCm.ai helps developers install, deploy, troubleshoot and optimize AI workloads on AMD platforms through a unified software experience.
Why It Matters? AI development is becoming increasingly agent-driven, with developers using natural language and AI coding assistants to build, deploy and optimize software. ROCm.ai brings AMD expertise directly into those workflows, helping developers build and optimize AI workloads using AMD-specific guidance instead of generic recommendations. The introduction of AMD Skills reduces barriers to adopting ROCm and reducing the amount of documentation crawling required for setup. ROCm.ai also uses AI to improve the AMD ROCm™ software stack itself. Through AI-driven optimization of kernels, memory management and scheduling, ROCm.ai delivers an average 3.3x inference improvement and 2.4x training improvement for the latest version of ROCm over ROCm™ 7 on the same hardware.1 2
What’s the Role for AMD?
AMD is committed to helping developers move faster than ever. ROCm.ai brings agentic AI development to AMD platforms – agents that don’t just answer questions, but profile, debug and drive workloads toward peak performance on AMD hardware. It’s the next evolution of the AMD AI software stack, shortening the distance between an idea and running workloads so developers build and iterate fast. Whether standing up an environment, serving a model or closing the gap to peak performance, ROCm.ai does the heavy lifting – right inside the coding assistants developers already use.
What is ROCm.ai? ROCm.ai is the AMD AI-native developer experience designed to help build, deploy and optimize faster across AMD platforms. It combines ROCm CLI, AMD Skills and AI-powered software optimization into a unified experience that simplifies AI development from installation through deployment, debugging and performance optimization.
ROCm CLI provides a unified command-line experience for installing, validating, serving, updating and troubleshooting AI workloads on AMD platforms. Hardware-aware workflows, automated environment setup and support for secure, air-gapped deployments help simplify software installation while enabling consistent deployment across environments.
AMD Skills integrates AMD-authored expertise into leading AI coding assistants including Claude, Cursor and Codex. Instead of relying on generic recommendations, developers receive AMD-specific guidance for installing, migrating, debugging and optimizing AI workloads across the AMD AI software stack.
Hyperloom automates the optimization of end-to-end inference workloads. It handles the full inference optimization loop, from profiling and analysis to kernel optimization and validation, helping accomplish in hours what previously required weeks of specialized engineering effort.
ith AI-powered software optimization, including Hyperloom, ROCm.ai helps improve the ROCm software stack across parallelization and scheduling, memory management and optimized kernels. These software improvements help accelerate inference and training performance on existing AMD hardware while continuing to improve developer productivity and workload performance over time.
Availability for ROCm.ai will begin in August 2026.
More: Advancing AI 2026 (Press Kit)
Cautionary Statement
This blog may contain forward-looking statements concerning Advanced Micro Devices, Inc. (AMD), which are made pursuant to the Safe Harbor provisions of the Private Securities Litigation Reform Act of 1995. Forward-looking statements are commonly identified by words such as "would," "may," "expects," "believes," "plans," "intends," "projects" and other terms with similar meaning. Investors are cautioned that any forward-looking statements in this blog are based on current beliefs, assumptions and expectations, speak only as of the date of this blog and involve risks and uncertainties that could cause actual results to differ materially from current expectations. Such statements are subject to certain known and unknown risks and uncertainties, many of which are difficult to predict and generally beyond AMD's control, that could cause actual results and other future events to differ materially from those expressed in, or implied or projected by, the forward-looking information and statements. Investors are urged to review in detail the risks and uncertainties in AMD’s Securities and Exchange Commission filings, including but not limited to AMD’s most recent reports on Forms 10-K and 10-Q.
AMD does not assume, and hereby disclaims, any obligation to update forward-looking statements made in this blog, except as may be required by law.
-
(MI350-81) Testing by AMD Performance Labs as of July 7, 2026, measuring the inference performance in tokens per second (TPS) of a system configured with an AMD Instinct MI355x 8x GPU platform and AMD ROCm 7.0 software vs a similarly configured system using a preview version of AMD ROCm.ai (ROCm 7.2.2 with optimizations such as Optimized Kernels, Parallelism and Scheduling) running GLM-5, Kimi-K2.5, and DeepSeekk-R1-0528 models.
Stated performance uplift is expressed as a combined average TPS over across the (3) models tested.
Hardware Configuration
Supermicro AS -4126GS-NMR-LCC (board H14DSG-OD)
8x AMD Instinct MI355X. BIOS AMI v1.4a (2025-04-16), GPU firmware SMC 04.86.11.02, TA RAS 27.69.00.10, TA XGMI 32.00.00.20, RLC43, MEC36, SDMA12, Ubuntu 22.04.2 LTS, kernel 5.15.0-70-generic, amdgpu driver 6.16.6, HOST ROCm 7.1.0
Software Configuration(s)
GLM-5 ROCm Docker Image: rocm/sgl-dev:v0.5.8.post1-rocm700-mi35x-20260219
PYTorch Version 2.8.0, SGLang v0.5.8
Kimi ROCm Docker Image: vllm/vllm-openai-rocm:v0.16.0, vLLM version 0.16.0
DeepSeek-R1 Docker Image: rocm/7.0:...sgl-dev-v0.5.2-rocm7.0-mi35x-20250915, SGLang version V0.5.13
vs
GLM-5 ROCm Docker Image: rocm/atom:rocm7.2.2_ubuntu24.04_py3.12_pytorch_release_2.10.0_atom0.1.2.post, ATOM v0.1.2.post
Kimi ROCm Docker Image: vllm/vllm-openai-rocm:v0.22.0, vLLM version V0.22.0
DeepSeek-R1 Docker Image: lmsysorg/sglang-rocm:v0.5.13-rocm720-mi35x-20260612, SGLang version V0.5.13
Server manufacturers may vary configurations, yielding different results. Performance may vary based on configuration, software, vLLM version, and the use of the latest drivers and optimizations. (MI350-81)
-
(MI350-82) Testing by AMD Performance Labs as of July 7, 2026, measuring the training performance in tokens per second (TPS) of AMD ROCm 7.0 software vs a preview version of AMD ROCm.ai (ROCm 7.2.2 with optimizations such as Optimized Kernels, Parallelism and Scheduling),) using Megatron -LM on a system with 8x AMD Instinct MI355x 8x GPUs GPU platform running DeepSeek-V2-Lite, DeepSeek-V3-16B, and Qwen3-30B-A3B models.
Stated performance uplift is expressed as the combined average TPS over across the (3) models tested.
Hardware Configuration
Supermicro AS -4126GS-NMR-LCC (board H14DSG-OD)
8x AMD Instinct MI355. BIOS AMI v1.4a (2025-04-16), GPU firmware SMC 04.86.11.02, TA RAS 27.69.00.10, TA XGMI 32.00.00.20, RLC 43, MEC 36, SDMA 12, Ubuntu 22.04.2 LTS, kernel 5.15.0-70-generic, amdgpu driver 6.16.6, HOST ROCm 7.1.0
Software Configuration(s)
DeepSeek-V2-Lite, ROCm 7.2.1 + Primus v26.3
DeepSeek-V3-16B, ROCm 7.2.1 + Primus v26.3
Qwen3-30B-A3B, ROCm 7.2.1 + Primus v26.3
Server manufacturers may vary configurations, yielding different results. Performance may vary based on configuration, software, and the use of the latest drivers and optimizations. (MI350-82)
Press inquiries: corporate.pressinquiry@amd.com